Back

Genomics, Proteomics & Bioinformatics

Preprints posted in the last 7 days, ranked by how well they match Genomics, Proteomics & Bioinformatics's content profile, based on 188 papers previously published here. The average preprint has a 0.11% match score for this journal, so anything above that is already an above-average fit.

1
SALRR: Scalable Analysis of Long-Read RNA-Seq Enables Comprehensive Transcriptome Profiling in Human Brain

Kouam, C.; Mingle, J.; Alvarez Jerez, P.; Evans, A.; Moller, A.; Baker, B.; Weller, C.; Paquette, K.; Brooks, J.; Grant, S. M.; Ayuketah, A.; Meredith, M.; Palade, J.; Malik, L.; Hise, K.; Raphael Gibbs, J.; Anderson, J.; Ding, J.; Harbert, R.; Fu, Y.; Zheng, X.; Garcia-Ruiz, S.; Gustavsson, E. K.; Blauwendraat, C.; Ryten, M.; Sedlazeck, F.; Ferrucci, L.; Reed, X.; Nalls, M. A.; Cookson, M. R.; Van Keuren-Jensen, K.; Hutchins, E.; Jain, M.; Billingsley, K. J.

2026-08-29 genomics 10.64898/2026.08.27.747499 medRxiv
Top 0.8%
2.6%
Show abstract

Isoform-resolved transcriptomics is fundamental to decoding the molecular complexity of the human brain, yet population-scale long-read RNA sequencing has remained inaccessible due to labor-intensive library preparation, sensitivity to RNA degradation in postmortem tissue, and the absence of integrated, reproducible analysis pipelines. Here we present SALRR (Scalable Analysis of Long-Read RNA-seq), an integrated wet-lab and computational platform designed to overcome these barriers. Automated ONT long-read cDNA library preparation on the Hamilton Microlab NGS STAR platform reduces hands-on time by 67% and enables 24 libraries per operator per day while maintaining performance across RNA integrity values. A modular, Snakemake-based pipeline performs end-to-end processing from ONT signal data to isoform-level quantification, incorporating SIRV spike-in calibration, multi-stage quality control, and stringent isoform validation. Applied to 10 postmortem frontal cortex samples from the North American Brain Expression Consortium, SALRR identified 31,607 high-confidence isoforms from 10,075 genes, including 8,532 novel splice variants absent from GENCODE v49, and complex splicing events systematically missed by short-read sequencing at neurodegeneration-relevant loci, including GBA1, CCNF, CHCHD10, and TREM2. All protocols and code are openly available, providing a scalable, community-ready framework for isoform-resolved transcriptomics in neurodegeneration, aging, and complex brain disease.

2
Reference-guided comparative genomics of seven Indonesian rice cultivars identifies conserved gene space and trait-associated sequence candidates

Purwestri, Y. A.; Wicaksono, A.; Nurbaiti, S.; Purba, N. T.; Retnaningati, D.; Restiani, R.; Kumalasari, N.; Nuringtyas, T. R.; Handayani, V. D. S.

2026-08-29 genomics 10.64898/2026.08.26.747264 medRxiv
Top 1%
1.9%
Show abstract

Indonesian rice cultivars represent valuable genetic resources, yet many remain poorly characterized at the genomic level. Here, we generated 95.40 Gb of PacBio HiFi sequence data from seven Indonesian rice cultivars and constructed cultivar-specific consensus genomes using the telomere-to-telomere Nipponbare reference AGIS1.0. Sequencing coverage ranged from 27.92x to 41.58x, and the resulting consensus genomes spanned 387.93-390.54 Mb, with BUSCO completeness of approximately 98.3-98.5%. OrthoFinder assigned 99.1% of predicted proteins to 40,737 orthogroups, including 27,514 core orthogroups represented across all seven cultivars, indicating a highly conserved predicted gene space within the reference-guided framework. Targeted analysis recovered 278 of 280 cultivar-by-locus combinations representing 40 genes or gene family entries associated with grain pigmentation, nitrogen and amino-acid metabolism, and starch properties. Comparative predicted protein analysis prioritized ANS1, SBE2b, SSIIa/ALK, Wx/GBSSI, OsAAP6/qPC1, and SSI as candidates for further investigation. Among 269 completed AGIS1.0-anchored promoter comparisons, 159 passed quality-control criteria, whereas 110 were flagged for gene-model, boundary, synteny, or structural concerns. Notably, these flagged comparisons accounted for more than 90% of the alignment-derived sequence variation, emphasizing the importance of rigorous quality control when interpreting apparent promoter divergence. Collectively, these reference-guided genomic resources provide a standardized framework for investigating sequence variation in Indonesian rice germplasm and prioritize testable coding and regulatory candidates for functional validation and future genomics-assisted crop improvement.

3
Decoding the Transcriptome Dark Matter: Construction of Single-Cell Whole-Transcriptome Regulatory Atlas by dropTotal

Liu, X.; Cao, W.; Pan, Y.; Luo, Z.; Wu, T.; Du, Y.; Xu, X.; Jin, Z.; Li, C.; Mu, Y.; Liu, Y.; Zhu, Q.

2026-08-29 genomics 10.64898/2026.08.25.747146 medRxiv
Top 1%
1.7%
Show abstract

To profile unknown ncRNAs-"dark matter" in single cells, we developed dropTotal, a high-throughput droplet-based total RNA-seq method that uses dU-modified GAT primer with temperature-ramp hybridization and droplet merge barcoding to co-detect coding and non-coding transcripts with record sensitivity (>13,500 genes/cell, including >2,000 lncRNAs and >500 sncRNAs), compatible with fresh, frozen, fixed, and FFPE tissues. Applied to ~75,000 human glioma nuclei, it captured 60,313 genes (18,681 lncRNA, 19,859 mRNAs and 6,753 sncRNAs), enabling ncRNA-driven regulatory landscape construction. In oligodendroglioma, module analysis identified recurrence-associated ncRNA-centered modules linked to therapy resistance and invasion; in glioblastoma, six cellular states showed hundreds of state-specific unannotated ncRNAs with divergent functions, from MIR222HG-mediated immune modulation to SCIRT-driven hypoxia adaptation. Alternative splicing analysis identified 428 state-specific junction markers and mapped cell-state-specific alternative splicing regulation. dropTotal offers broad application for decoding the underlying ncRNA biology and single-cell whole transcriptome regulatory mechanisms in cellular identity and disease progression.

4
Perturb-seq identifies co-regulated gene programs shaping hematopoietic stem and progenitor cell function

Bowness, J. S.; Bernal Martinez, A.; Barinka, J.; Schulte-Schrepping, J.; Renders, S.; Waclawiczek, A.; Leppa, A.-M.; Trumpp, A.; Raffel, S.; Haas, S.; Velten, L.

2026-08-29 genomics 10.64898/2026.08.27.747033 medRxiv
Top 2%
1.0%
Show abstract

To sustain blood formation, hematopoietic stem and progenitor cells (HSPCs) coordinate a multitude of cell biological processes, from cell cycle control and stress responses to lineage priming. While many genetic regulators of high-level HSPC function have been identified, how HSPCs coordinate more basal cell biological programs, and how such programs relate to stem cell function, remains incompletely understood. Here we use Perturb-seq to profile the transcriptional consequences of targeting 520 genes by CRISPRi in primary mouse HSPC cultures. We developed an analytical strategy to separate perturbation-induced changes in cell-state abundance and clonal heterogeneity from cell-state-local transcriptional effects. From these local perturbation signatures, we identified 19 gene regulatory programs (GRPs) that are defined by co-regulation in response to genetic perturbation, in contrast to co-expression or human curation, and align well with cell biological processes. By decomposing gene expression data from functional and clinical studies into program activity, we show that GRP activities associate with, and predict, phenotypes such as clonal output after transplantation, as well as survival and drug response in retrospective acute myeloid leukemia (AML) cohorts. Together, our study establishes perturbation-derived co-regulation programs as an interpretable framework for linking genetic regulators, cell-biological processes and stem-cell-associated phenotypes.

5
Pan-cancer Graph-based Cancer Detection Using the Cell-free DNA Methylome

Zhao, L.; Zeng, Y.; Abelman, D. D.; Lin, W.; Luo, P.

2026-08-31 oncology 10.64898/2026.08.26.26361432 medRxiv
Top 2%
1.0%
Show abstract

Motivation: Cell-free DNA methylation provides a minimally invasive signal for early cancer detection and tissue-of-origin prediction. Most methods represent methylation measurements as independent fixed-window features and therefore do not explicitly model relationships among genomic regions. Results: We developed PANGEM (Pan-cancer Graph-based Cancer Detection Using the Cell-free DNA Methylome), a graph-learning framework that represents genomic bins as nodes and integrates CpG context, genomic proximity, and sample-specific methylation similarity in the graph topology. Across five repeated stratified train-test splits, PANGEM achieved the highest mean performance among evaluated methods, with an AUROC/AUPR of 0.997/1.000 for binary cancer detection and macro-AUROC/AUPR of 0.977/0.870 for multiclass tissue-of-origin prediction. In the independent INSPIRE cohort, 72 of 78 cancer cases (92.3%) exceeded the binary classification threshold, and PANGEM correctly classified 9 of 17 head and neck cancer cases (52.9%), the highest accuracy among evaluated methods. Subnetwork analysis further identified recurrent, graph-connected methylation patterns, including a 111-DMR subnetwork with increased methylation in cancer samples.

6
MOSurvivor-Guided Joint CpG Selection and XGBoost Hyperparameter Optimization for Compact Epigenetic Age Prediction

Yelgi, A.; Tavangari, S.; Shakarami, Z.; Janfaza, S.

2026-08-29 genomics 10.64898/2026.08.26.747213 medRxiv
Top 2%
0.9%
Show abstract

Accurate epigenetic age prediction from DNA methylation profiles is intrinsically high-dimensional, creating a need for parsimonious models that preserve predictive performance while reducing the number of assayed cytosine-phosphate-guanine (CpG) loci. This study introduces MOSurvivor, a population-based multi-objective search framework that jointly optimizes a weight-threshold CpG selector and eight XGBoost hyperparameters. Experiments used the GSE40279 whole-blood cohort (656 individuals profiled on the Illumina HumanMethylation450 platform). After retaining 1,000 age-correlated CpGs, five strategies were evaluated on the same 30 seeded 80:20 train/test splits: fixed-parameter XGBoost using all 1,000 CpGs, random search, a genetic algorithm, particle swarm optimization, and MOSurvivor. Internal fitness was estimated using three-fold cross-validation on each training set. Across the 30 held-out test sets, MOSurvivor achieved a mean absolute error (MAE) of 4.149 {+/-} 0.300 years, root mean squared error of 5.545 {+/-} 0.392 years, and R2 of 0.855{+/-} 0.027 while retaining 211.6 {+/-} 54.8 CpGs. Relative to full-feature XGBoost (MAE 4.095 {+/-} 0.285 years), MOSurvivor reduced the feature set by 78.8% at an MAE increase of only 0.054 years (1.3%). Paired Wilcoxon tests found no significant accuracy difference between MOSurvivor and any comparator (all unadjusted p > 0.05; all Holm-adjusted p [≥] 0.476). The most recurrent locus, cg16867657, appeared in 29 runs, whereas mean pairwise Jaccard similarity was 0.124, indicating a small stable core embedded in multiple near-equivalent feature subsets. MOSurvivor thus offers a competitive accuracy-parsimony trade-off rather than superior absolute accuracy. External validation and leakage-free nested feature preselection remain necessary before biological or clinical translation. Keywords: epigenetic clock, DNA methylation, feature selection, multi-objective optimization, XGBoost, metaheuristics, biological aging.

7
A Simple Method to Distinguish Active and Inactive Aptamers by Analyzing the Ruggedness of the Aptamer Free Energy Landscape

Subramanian, G.; Thiel, W.; Singh, R.

2026-08-29 bioinformatics 10.64898/2026.08.26.747184 medRxiv
Top 2%
0.8%
Show abstract

Aptamers are structured nucleic acid ligands capable of high affinity, high specificity molecular recognition generated using variations of the SELEX (Systematic Evolution of Ligands by Exponential Enrichment) process. However, SELEX often produces sequences that enrich yet may lack binding efficacy. We propose a measure called the Ruggedness Composite Index (RCI) along with a method for computing it, that can be used to distinguish binding-competent ('active') aptamers from weak or non-binding ('inactive') aptamers. Given a set of aptamers, RCI incorporates information on their fragmentation (landscape partitioning), basin entropy (metastable state distribution), cumulative density irregularity (non-uniform occupancy), and structural energy correlation length (structure-energy coupling scale). We test whether secondary-structure folding energy landscape topology distinguishes active from inactive aptamers using a multiscale level set framework across six datasets. Active aptamers show lower RCI values and occupy smoother, funnel-like conformational spaces, while inactive aptamers show higher RCI values, reflecting fragmented, high-entropy landscapes. By contrast, classical thermodynamic features, such as minimum free energy, show limited discrimination between active and inactive aptamers. In all datasets, sequences that exhibit enrichment which is not monotonic but lack specificity exhibit elevated ruggedness, indicating landscape topology can predict non-specific enrichment. These results indicate that folding landscape organization can be used as a predictor of aptamer activity and establish RCI as a simple, mechanistically interpretable measure for improving candidate prioritization, especially in therapeutic aptamer discovery.

8
Visual LLM-guided consensus spatial domain detection with L-STAR

Zhao, C.; Ji, Z.

2026-08-29 bioinformatics 10.64898/2026.08.25.747158 medRxiv
Top 3%
0.6%
Show abstract

Spatial domain detection is a central task in spatial transcriptomics, yet existing methods exhibit highly variable performance across datasets. We introduce L-STAR, a visual LLM-guided, consensus-based framework that leverages the visual reasoning capacity of large language models to adaptively rank and integrate spatial domain detection methods. L-STAR achieves robust and consistently improved performance, outperforming single spatial domain detection methods across diverse datasets.

9
Persistence of Extended Spectrum β-Lactamase-Producing Enterobacterales in the Gut Microbiome of Healthy Newborns

Shuai, W.; Mithal, L. B.; Kremer, A.; Aron, A.; Sajwani, A.; Huntinghouse, D.; Hartmann, E. M.; Arshad, M.

2026-09-03 infectious diseases 10.64898/2026.09.01.26361559 medRxiv
Top 3%
0.5%
Show abstract

The global prevalence of Extended-spectrum {beta}-lactamase-producing Enterobacterales (ESBL-E) colonization is increasing. However, it is unclear whether ESBL-E persist and if that is associated with an altered gut microbial ecology especially in early life where the developing microbiome may not provide the same colonization resistance as in adults. In this study, we collected longitudinal infant gut microbiome samples at delivery and in the nonclinical home setting in Chicago, Illinois, U.S.A, aiming to disentangle how genetic factors pertaining to the ESBL-E, as well as the surrounding gut ecology, influences persistence in the infant gut microbiome. We observed not only a higher-than-expected prevalence of ESBL-E in healthy infant gut microbiomes, but also a trend of ESBL-E persistence once colonized. Microbial communities showed higher dissimilarity between ESBL-E positive and negative infant gut microbiome at earlier time points. Although dissimilarity decreased over time, we present evidence that ESBL-E persist even when traditional detection methods are negative.

10
The QxxR Motif of RNA Helicase Me31B Is Essential for Drosophila Female Fertility and Germline Development

Mansoor, R.; Minhas, A. S.; Thomas, A.; Mansoor, A. A.; McCambridge, A. H.; Dilts, C.; Eshak, J.; Govani, D.; Nylin, B.; Trinidad, J. C.; Kanaan, A. Y.; Kara, E.; Fielder, A.; Fielder, I.; Iglendza, A.; Mukatash, Y.; Pumnea, B.; Menzel, M. M.; Shabazz-Henry, A. L.; Niepielko, M. G.; Gao, M.

2026-08-29 genetics 10.64898/2026.08.27.747641 medRxiv
Top 3%
0.4%
Show abstract

The QxxR motif is evolutionarily conserved within DEAD-box RNA helicases, including Drosophila Me31B and human DDX6, which post-transcriptionally regulate gene expression during animal development. A pathogenic H372R substitution (QxHR to QxRR) in the QxxR motif of human DDX6 has been associated with various developmental defects, but how this motif contributes to DDX6-family protein function remains unclear. Here, we used Drosophila Me31B as an in vivo model to investigate the QxxR motifs developmental role. We generated a Drosophila strain carrying the corresponding H333R missense mutation in Me31B and characterized its effects on female fertility, embryonic viability, germline development, and Me31B-associated molecular pathways. The me31BH333R mutation reduced female fertility in a gene dose-dependent manner, with homozygous mutant females being sterile. Embryos from the mutant females also exhibited primordial germ cell defects. Despite these developmental phenotypes, the me31BH333R mutation did not significantly alter Me31B protein abundance, global ovarian transcriptome or proteome profiles, or representative germ plasm mRNA and protein localization. In contrast, bait-normalized IP-MS analysis revealed altered enrichment of selected Me31B-associated proteins, including increased association of known Me31B interactors Trailer hitch (Tral) and Ypsilon Schachtel (Yps). These findings establish Me31BH333R as an in vivo model for investigating the conserved QxxR motif and suggest that disruption of this motif compromises development not through broad changes in gene expression, but potentially through altered composition or regulation of Me31B-containing ribonucleoprotein complexes.

11
Utilising nuclear encoded plastid DNA to identify donors of grass-to-grass lateral gene transfer

Bourne, N. G.; Payne, L.; Manzi, S.; Besnard, G.; Vorontsova, M. S.; Jobson, R. W.; Chomicki, G. S.; Dunning, L. T.

2026-08-29 evolutionary biology 10.64898/2026.08.26.747220 medRxiv
Top 5%
0.3%
Show abstract

Determining the correct donor species/lineages of grass-to-grass lateral gene transfer (LGT) is vital for deducing specific donor features that could help inform the mechanism of transfer. This requires a dataset spanning a broad range of species to achieve the phylogenetic resolution necessary for precise donor inference. As grass-to-grass LGT often involves the transfer of multi-gene DNA fragments, they can contain additional sequences that allow for accurate orthologous comparisons, such as nuclear DNA of plastid origin (NUPTs). Here we systematically scan for NUPTs in the genomes of four Alloteropsis semialata accessions, whose LGTs have previously been characterised. Using the abundant Panicoideae chloroplast sequences, we reconstruct NUPT phylogenies and infer two lateral acquisitions: one from Paniceae/Digitaria and another from Andropogoneae/Eremochloa adjacent to a previously identified LGT. We then assembled and included an additional 12 Eremochloa chloroplast genomes in the analysis and showed the likely donor was Eremochloa attenuata. Subsequent short-read mapping from E. attenuata to the nuclear region flanking this NUPT showed consistent coverage across the region, including the previously identified LGT, supporting co-transfer. Overall this study highlights the potential for NUPTs to better identify the donors of grass-to-grass LGT.

12
Quantifying the Rearrangement Complexity of Pangenomes

Bohnenkaemper, L.; Stoye, J.

2026-08-29 bioinformatics 10.64898/2026.08.27.747493 medRxiv
Top 5%
0.3%
Show abstract

The study of evolution between species (phylogenetics) and the study of evolution within a species (population genetics) are highly related, as the same biological mechanisms are fundamental to both fields. Although both have been studied for a long time, their joint study in a unified setting has been prevented by the different time scales they consider and the different data types they employ. A similar discrepancy holds for their whole-genome specializations, comparative genomics and pangenomics. Two active areas in these fields are genome rearrangement studies and graphical pangenomics, respectively. Since the emergence of graphical pangenomics, these have existed as separate fields, despite observations that central data structures representing genomic variants in both fields are highly similar. While there exists a wealth of theoretical results for various rearrangement models in comparative genomics, the application to pangenomic data is hampered by the limitations of rearrangement problem formulations. On the practical side, pangenomes typically contain too many individual genomes for classical problems, such as the often NP-hard parsimony problems, to be solved, or for all-vs-all comparisons using rearrangement distances to be performed. On the theoretical side, some assumptions in the formulation of rearrangement problems, such as the assumption of an underlying tree, are inadequate for many pangenomes. In this work, we propose the Complete Ancestral Reconstruction for Pangenomes (CARP) problem, which overcomes these limitations while retaining intuitive relationships to both classical rearrangement problems and pangenome graphs.

13
Cost-Utility Analysis of First-Line Olaparib plus Abiraterone for Metastatic Castration-Resistant Prostate Cancer in China after Volume-Based Procurement

SHI, J.; Gu, Q.; Pan, J.; Yang, A.; Fan, M.

2026-08-31 health economics 10.64898/2026.08.26.26361314 medRxiv
Top 5%
0.2%
Show abstract

To evaluate the cost-utility and 5-year budget impact of first-line olaparib plus abiraterone versus abiraterone alone for metastatic castration-resistant prostate cancer (mCRPC) in China after the eleventh round of volume-based procurement (VBP). The intention-to-treat (ITT) population was assigned primary decision-analytic weight; the prespecified BRCA1/2-mutated (BRCAm) subgroup was a supporting analysis.

14
Bacterial metagenome in plaque, saliva, and tumor samples from individuals with and without OSCC by next-generation sequencing

ERIRA, A.; ROBAYO, D. A. G.; GAMBOA, F.; CHALA, A.; MORENO, A.; ARREGUI, A. C.; MUNOZ, E.; NOGUERA, J.; TOBAR-TOSSE, F.

2026-08-29 bioinformatics 10.64898/2026.08.27.747557 medRxiv
Top 5%
0.2%
Show abstract

Background: Oral dysbiosis has been associated with oral squamous cell carcinoma (OSCC); however, most microbiome studies rely on 16S ribosomal RNA (rRNA) gene sequencing, limiting species-level taxonomic resolution. Methods: Dental plaque, saliva, and tumor tissue samples from 10 patients with OSCC and dental plaque and saliva samples from 10 healthy controls were analyzed in this exploratory cross-sectional study. DNA was extracted and subjected to shotgun metagenomic sequencing using the Illumina MiSeq platform. Sequence reads were quality filtered with fastp, taxonomically classified using Kraken2 v2.1.3, and species-level abundances were re-estimated with Bracken v2.9 following the removal of human reads and low abundance taxa. Relative abundances were compared using the Mann Whitney U test with the Benjamini Hochberg false discovery rate correction, while the Bray Curtis principal coordinate analysis was used as an exploratory approach to visualize microbial community patterns. Results: Shotgun metagenomic sequencing revealed distinct bacterial community profiles across the oral microenvironment. Dental plaque exhibited the highest taxonomic diversity and relative abundance. The control plaque was enriched in Streptococcus koreensis, Capnocytophaga sp. oral taxon 878, Treponema sp. Marseille Q4132, and Leptotrichia sp. oral taxon 498, whereas the plaque from patients with OSCC showed a higher relative abundance of Pyramidobacter piscolens, Parvimonas parva, and Gemella sanguinis. Salivary samples displayed lower diversity and a more homogeneous composition, predominantly comprising Capnocytophaga endodontalis, Prevotella jejuni, Aggregatibacter aphrophilus, and Gemella sanguinis. The tumor tissue showed relatively higher abundance of Sellimonas catena, Escherichia coli, Solobacterium moorei, and Lacrimispora sp. HJ 01. Conclusions: This exploratory study provides species-level characterization of the oral microbiome across multiple oral microenvironments in OSCC and generates hypotheses for future integrative metagenomic and functional studies investigating the potential contribution of oral bacterial communities to OSCC pathogenesis.

15
Molecular landscape and risk stratification in acute myeloid leukemia - insights from the real-world REFORM-AML cohort

Kristensen, D. T.; Broendum, R. F.; Knudsen, M.; Grubach, L.; Marcher, C.; Preiss, B.; Bibi, M. L.; Hoegdall, E.; Poulsen, T.; Skov, V.; Oerskov, A. D.; Groenbaek, K.; Hansen, J. W.; Schoellkopf, C.; Cowland, J.; Andersen, M. K.; Severinsen, M. T.; Vejgaard, C.; Larsen, O. H.; Vang, S.; Boegsted, M.; Roug, A. S.

2026-08-31 hematology 10.64898/2026.08.27.26361552 medRxiv
Top 5%
0.2%
Show abstract

Large genomically annotated acute myeloid leukaemia (AML) datasets exist, but population-based contemporary cohorts remain scarce. Here we report clinicopathological, genomic, and outcome data from Danish AML patients. 2,512 AML patients were identified between 2015-2022, of whom 33.8% had available NGS data (NGS+). In patients [≤]70 years, baseline characteristics and outcomes were comparable between NGS+ and NGS- groups. In patients >70 years, more NGS+ patients received intensive treatment, but survival was similar among intensively treated patients. The distribution of mutations varied significantly by age and sex, with older age and male sex exhibiting higher frequencies of adverse-risk gene mutations. In intensively treated NGS+ patients, ELN2017 stratified 5-year OS: 58.4% (favorable), 43.4% (intermediate), and 28.2% (adverse), with hazard ratios (HRs) of 0.63 (favorable) and 1.45 (adverse) relative to intermediate. ELN2022 yielded corresponding OS rates of 56.9%, 51.8%, and 29.7%, with HRs of 0.78 and 1.86. The two models had comparable predictive performance for OS in a time-dependent model. In conclusion, outcomes of intensively treated AML patients were comparable irrespective of NGS status, underscoring the representativeness of the REFORM-AML database for the Danish AML population. Age and male sex correlated with adverse-risk mutations, and both ELN2017 and ELN2022 robustly predicted survival.

16
Geometric characterization of the HSV - 1 glycoprotein B - amyloid β interaction in Alzheimer's disease using Forman-Ricci curvature

Bou Dagher, L.; Han, Z.; Zhou, S.; Fülöp, T.; Desroches, M.; Rodrigues, S.

2026-08-29 bioinformatics 10.64898/2026.08.26.747308 medRxiv
Top 5%
0.2%
Show abstract

Alzheimer's disease is characterized by the accumulation and aggregation of amyloid-{beta}(A{beta}), but the molecular mechanisms linking environmental and infectious factors to A$\beta$ conformational changes remain incompletely understood. Herpes simplex virus type 1 (HSV-1) has been proposed as a potential contributor to AD pathology, and interactions between the viral glycoprotein B (gB) and A$\beta$ may influence the conformational behaviour of the peptide. Molecular dynamics (MD) simulations provide atomic-scale information on such interactions, but conventional structural descriptors may not fully capture changes in the organization of residue interaction networks. Here, we introduce a graph-geometric framework based on Forman-Ricci curvature to characterize the evolution of residue interaction networks during MD simulations. Each simulation frame is represented as a residue interaction graph based on C--C contacts, and residue-wise curvature profiles are analysed across time. We apply the framework to A{beta}1-42 in isolation and in complex with HSV-1 gB. Conventional MD analyses indicate stable association of the simulated complex, favourable interaction energetics, and conformational changes in A{beta}, including a transition from -helical structure toward {beta}-turn-rich conformations over the simulated timescale. Forman-Ricci curvature reveals pronounced and spatially localized remodelling of the A{beta} residue interaction network in the complex, with the strongest changes concentrated in the C-terminal region. These regions also exhibit reduced temporal curvature fluctuations and progressively distinct geometric behaviour throughout the simulation. Hierarchical clustering further identifies cooperative groups of residues with coordinated curvature dynamics, including a prominent C-terminal domain. Together, these results demonstrate that Forman-Ricci curvature provides a complementary description of biomolecular dynamics by capturing changes in the geometric organization of residue interaction networks that are not directly represented by conventional structural descriptors. The framework provides a general computational approach for studying network-level structural remodelling in protein molecular dynamics and offers a quantitative perspective on the conformational consequences of HSV-1 gB--A{beta} association.

17
Benchmarking ten frontier large language models on 1,477 board style multiple choice questions in hematology

Radoynova, M.; Benouis, M.; schulze, f.; Winter, S.; Bornhauser, M.; Middeke, J. M.; Eckardt, J.-N.

2026-09-02 hematology 10.64898/2026.09.01.26361881 medRxiv
Top 6%
0.2%
Show abstract

Large Language Models (LLMs) are increasingly used by clinicians and patients for medical queries, yet their accuracy and safety at the specialist level in hematology remain insufficiently characterised. We benchmarked ten frontier proprietary and open-weight LLMs across two generations on 1,477 board-style hematology multiple-choice questions (MCQs) derived from five educational datasets spanning nine disease areas and six clinical skill domains, including text-only and multimodal case vignettes. Claude Opus 5 had the highest mean accuracy (92.7% text, 76.9% multimodal), followed closely by Gemini-3.1 Pro (91.4% and 78.7%), Gemini-3.6 Flash (91.0% and 74.8%) and GPT-5.6 Sol (89.9% and 76.7%). Accuracy significantly correlated with model size both for text-only and multimodal MCQs. Between model generations, the largest improvements in accuracy were seen for open-weight models whereas proprietary models showed only marginal gains. In error analysis, top-performing models exhibited highly concordant failure patterns, suggesting shared limitations on challenging cases. Frontier LLMs exhibit substantial specialist hematology knowledge across diverse subspecialist domains and clinical skill sets. Yet, despite high accuracy on board-style questions in hematology, continuous expert-on-the-loop output monitoring is paramount.

18
GLP-1/GIP Uptake, Indication, and Access Pathways Among US Adults in the Understanding America Study

Chaturvedi, R. R.; Gracner, T.; Perez-Arce, F.; Suen, S.-c.; Jin, J.; Orriens, B.; Pacula, R. L.; Sexton Ward, A.; Haile, R.; Kapteyn, A.

2026-09-02 endocrinology 10.64898/2026.08.28.26361368 medRxiv
Top 6%
0.1%
Show abstract

Importance: Evidence on GLP-1/GIP therapies is largely derived from trials enrolling selected populations or medical records that miss utilization outside healthcare channels. No nationally representative cohort has characterized real-world uptake, indications, and access. Objective: To characterize GLP-1/GIP prevalence, indication, clinical profile, and access. Design: Prospective cohort study with three GLP-1/GIP surveillance waves (March 2024, December 2024, October 2025). Setting: The Understanding America Study, an address-based, nationally representative panel of approximately 15,000 US adults aged 18+ years initiated in 2014. Participants: UAS participants responding to at least one surveillance wave (n=9150). Exposures: GLP-1/GIP use status (never vs any use, comprising current and former use), self-reported primary indication (diabetes, weight loss, or other), and access pathway (traditional vs non-traditional). Main Outcomes and Measures: Survey-weighted prevalence of GLP-1/GIP use, overall and by indication and access pathway; sociodemographic, cardiometabolic, treatment, and access characteristics; and smartwatch-derived resting heart rate, heart rate variability, maximum activity heart rate, step count, and sleep duration and variability. Results: Among n=9150 adults (1274 with any use; 60.9% female; median age 53 years), weighted prevalence increased 46%, from 8.2% (March 2024) to 12.0% (October 2025) representing 32 million. Weight-loss indications grew, reaching nearly half of use (4.1% to 5.6%); diabetes-indicated use was stable (5.3% to 5.4%). Users carried high cardiometabolic burden (obesity, 68.2%; diabetes, 53.6%) but diverged by indication: diabetes-indicated users were older (median, 59 vs 49 years), whereas weight-loss-indicated users were more often female (69.9% vs 51.3%) and healthier. One in three users (~9 million) had non-traditional access, especially in weight-loss-indicated users, of whom 33% had no conventional prescription; 41% used compounding, online, or foreign pharmacies; and, 43% lacked coverage. Non-traditional users were five times as likely to report an unlisted, likely compounded formulation (19.8% vs 4.1%). All p<0.05. Conclusions and Relevance: Real-world GLP-1/GIP use has grown rapidly and diversified substantially in indication, access, and population profile. One in 3 users obtained treatment through nontraditional channels largely invisible to claims data, raising long-term safety, efficacy, and coverage questions. GLIMMER provides a public, nationally representative longitudinal evidence base for future payer and provider decisions.

19
Can Dental AI Really Beat Dentists? DentalPair-Cert for Rigorous AI-Dentist Inference

Alve, S. R.; Rahman, S.; Meem, S. M. A. C.

2026-09-02 dentistry and oral medicine 10.64898/2026.09.01.26361874 medRxiv
Top 7%
0.1%
Show abstract

A dental AI system and a dentist reading the same radiographs form a paired comparison. Published comparative studies often report the two arms separately against a reference standard, leaving the joint pattern of correctness between them unavailable for secondary paired inference. We show what that omission costs. The accuracy difference remains exactly identified; its sampling variance does not, so the report contains the estimate and not its uncertainty. On a study of 282 units, two published accuracies are consistent with 38 distinct joint tables whose confidence intervals differ in width by a factor of 2.5. The consequence is a three-zone decision map rather than a single threshold: differences at or below 1.06 points are non-significant under every compatible table, differences at or above 6.03 points are significant under every compatible table, and in between the published numbers cannot decide. We then show the omission is repairable at negligible cost. One additional integer, the number of units both arms classify correctly, identifies the joint table exactly and restores standard paired inference. For a panel of readers the pairwise dependences must arise from one joint distribution, a constraint that binds once three readers are present; publishing each reader's joint-correct count against a single reference reader cannot widen and may tighten every pairwise bound, and in a 7-arm experiment reduced them by a median of 37% even for pairs excluding that reference. Where the integer was never published we give DentalPair-Cert, an interval with finite-sample coverage uniformly over every admissible within-unit AI-dentist dependence under the independent-sampling-unit model, certified in both the nuisance maximization and the inversion. Across 4,200,000 simulated comparisons an independence analysis falls to 74.5% coverage with 12.2% type-I error; in a purposive sample of 9 recent comparative studies, 1 reported a paired test on discordant units.

20
RedFuMOS: A novel approach for multi-omics and clinical data-driven patient stratification

De Luca, S.; Fava, C.; Rizzo, G.; Visconti, A.; Berchialla, P.

2026-08-31 health informatics 10.64898/2026.08.26.26361415 medRxiv
Top 7%
0.1%
Show abstract

Background. Patient stratification from multi-omics and clinical data is essential for uncovering disease heterogeneity and moving toward more personalized treatment strategies. However, integrating heterogeneous data layers while identifying robust patient strata remains challenging. Methods. We introduce Reduced Fusion of Multi-Omics Stratification (RedFuMOS), a novel three-step approach for patient stratification based on mixed-type multi-omics data. RedFuMOS extends Similarity Network Fusion to accommodate mixed-type data layers and layer-specific similarity measures for data integration, includes a dimensionality reduction step to mitigate the curse of dimensionality, and performs patient stratification using density-based hierarchical clustering with HDBSCAN. It also implemented an automated optimization procedure to identify the best set of hyperparameters, minimizing the need for manual tuning. Results. RedFuMOS outperformed six state-of-the-art tools for multi-omics patient stratification in a comprehensive simulated benchmarking study, which also confirmed that, although computationally expensive, the dimensionality reduction step is crucial for achieving good stratification performance. Additionally, RedFuMOS identified two clinically relevant patient strata in a small real-world cohort of patients with Philadelphia chromosome-positive chronic myeloid leukaemia. Conclusion. RedFuMOS provides a flexible framework for integrating heterogeneous multi-omics and clinical data. RedFuMOS is available as an R package at http://github.com/delucasara/RedFuMOS.